Papers with game design
Measuring Large Language Models’ Adversarial Behavior in Social Deduction Games (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing safety evaluations focus on refusal-based methods that test whether models avoid responding to inappropriate or violent requests, leaving open questions about how models behave in interactive social settings. |
| Approach: | They propose to use a meta-LLM to construct a closed behavioral taxonomy from a multi-agent simulation to examine adversarial behavior of large language models. |
| Outcome: | The proposed model-based model-driven model-model-based taxonomy shows that the model-led model-learning model exhibits distinct behavioral profiles and influences social stability and competitive success. |